Papers with computational analysis

14 papers
OpenFraming: Open-sourced Tool for Computational Framing Analysis of Multilingual Data (2021.emnlp-demo)

Copied to clipboard

Challenge: Existing frameworks for analyzing frames in multilingual text documents are available online and via an API.
Approach: They propose a web-based system for analyzing frames in multilingual text documents . framework combines unsupervised and supervised machine learning and leverages a state-of-the-art multilingual language model .
Outcome: The proposed framework can significantly improve frame prediction performance while requiring a small sample of manual annotations.
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are being used in urban planning but there is concern that they reproduce or amplify such biases.
Approach: They propose a framework to evaluate spatial gender bias in large language models . they use a taxonomy of 62 urban micro-spaces, a prompt library and three diagnostic layers .
Outcome: The proposed framework identifies structured gender-space associations that go beyond the public-private divide, forming nuanced micro-level mappings.
On the Impact of Temporal Representations on Metaphor Detection (2022.lrec-1)

Copied to clipboard

Challenge: State-of-the-art approaches for metaphor detection compare their literal - or core - meaning and their contextual meaning using neural networks.
Approach: They propose to use temporal and static word embeddings to account for different representations of literal meanings to examine metaphor detection tasks.
Outcome: The proposed method outperforms static methods but may provide representations of the core meaning of the metaphor too close to their contextual meaning, causing confusion.
The Timing of Lexical Memory Retrievals in Language Production (N18-1)

Copied to clipboard

Challenge: In a large-scale observational study of a spoken corpus, we find that language production at a time point preceding a word is sped up or slowed down depending on activation of that word.
Approach: They propose a cognitive model of fluency in which lexical memory retrievals may explain some of the variability in speech rates.
Outcome: The proposed model predicts that language production is sped up or slowed down depending on activation of a word .
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)

Copied to clipboard

Challenge: a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well.
Approach: They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data .
Outcome: The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP .
A Taxonomy of Empathetic Questions in Social Dialogs (2022.acl-long)

Copied to clipboard

Challenge: Current dialog generation approaches do not model effective question-asking due to the lack of a taxonomy of questions and their purpose in social chitchat.
Approach: They propose to model questions' ability to capture communicative acts and their emotion-regulation intents by annotating a large dataset with established labels.
Outcome: The proposed model can be used to generate labels for the EmpatheticDialogues dataset and to further improve the existing models.
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to analyze moral reasoning are discordant and lack cohesion, focusing on isolated aspects of the process.
Approach: They propose a unified dataset that integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, and captures diverse socio-cultural contexts.
Outcome: The proposed dataset integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, along with annotators’ moral and cultural profiles.
HisDoc-OCR: Restoring Visual Grounding in MLLMs for Chinese Historical Document OCR (2026.findings-acl)

Copied to clipboard

Challenge: Despite multimodal large language models' strong performance on modern document OCR, their application to historical Chinese texts suffers from severe hallucinations, character fabrication, uncontrolled repetition, and semantic drift.
Approach: They propose a multimodal large language model which restores visual grounding through three synergistic strategies: Layout Injection, First-Occurrence Boost, Self-Distilled Attention Focusing and HisDoc-OCR.
Outcome: The proposed model outperforms general-purpose and OCR-specific models on Chinese historical documents.
A Corpus of Natural Multimodal Spatial Scene Descriptions (L18-1)

Copied to clipboard

Challenge: Existing work on multimodal spatial descriptions combines speech and hand gestures to form a corpus of multimodal descriptions.
Approach: They present a corpus of multimodal spatial descriptions as commonly occurring in route giving tasks.
Outcome: The proposed corpus of multimodal spatial descriptions is more amenable to computational analysis and useable for learning natural computer interfaces.
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have an impressive ability to draw on novel information supplied in their context, yet the mechanisms underlying contextual grounding remain unknown.
Approach: They propose a method to study grounding abilities using a counterfactual dataset constructed to clash with a model's parametric knowledge using Fakepedia.
Outcome: The proposed method evaluates grounding abilities when the internal parametric knowledge clashes with the contextual information.
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic datasets embody a range of domain, dialect, and quality.
Approach: They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects.
Outcome: The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing.
Investigating Sports Commentator Bias within a Large Corpus of American Football Broadcasts (D19-1)

Copied to clipboard

Challenge: a recent study shows that sports broadcasters build drama into play-by-play commentary by building team and player narratives through subjective analyses and anecdotes.
Approach: They use FOOTBALL to examine racial bias in sports commentary . they identify major confounding factors for researchers examining rraecial bias .
Outcome: The proposed dataset supports previous social science studies on commentator bias . it contains 1,455 broadcast football transcripts annotated with 250K player mentions and racial metadata .
A Dual-View Analysis of Multiple Languages in Colonial Newspapers (2026.findings-acl)

Copied to clipboard

Challenge: Historical newspapers from the colonial period offer valuable evidence of how racializing language evolved over time.
Approach: They propose a contextual question answering and visual question answering task from colonial newspapers . they propose linguistic training for temporal word embedding with a compass to study racialization .
Outcome: The proposed tasks are limited for low-resource tasks, the authors show . the authors compare the results of two QA pairs from colonial newspapers to a compass .
N-CORE: N-View Consistency Regularization for Disentangled Representation Learning in Nonverbal Vocalizations (2025.emnlp-main)

Copied to clipboard

Challenge: Nonverbal vocalizations are an essential component of human communication, conveying rich information without linguistic content.
Approach: They propose a backbone-agnostic framework to disentangle emotion and speaker information from nonverbal vocalizations by leveraging N views of audio samples to learn invariance to specific transformations.
Outcome: The proposed framework achieves competitive performance compared to state-of-the-art methods on the VIVAE, ReCANVo, and ReCANVO-Balanced datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations